Papers with multimodal CoT module
M3Hop-CoT: Misogynous Meme Identification with Multimodal Multi-hop Chain-of-Thought (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have shown that Large Language Models (LLMs) neglect cultural diversity and key aspects like emotion and contextual knowledge hidden in the visual modalities. |
| Approach: | They propose a framework for misogynous meme identification using a multimodal multimodal prompting principle and a CLIP-based classifier. |
| Outcome: | The proposed framework performs well on the SemEval-2022 task 5 dataset, and is generalizable across different datasets. |